Back

Cancer Epidemiology, Biomarkers & Prevention

American Association for Cancer Research (AACR)

Preprints posted in the last 30 days, ranked by how well they match Cancer Epidemiology, Biomarkers & Prevention's content profile, based on 20 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Rigorous Female Breast Cancer Phenotyping Using the All of Us Research Program

Qi, Y.; Lundy-Perez, K.; Gee, D. A.; Chambwe, N.

2026-08-10 oncology 10.64898/2026.08.07.26359972 medRxiv
Top 0.1%
18.8%
Show abstract

Objectives Accurate phenotyping of cases and controls is essential for studying biological and environmental contributors to disease in large biobanks. We aimed to develop a flexible, customizable, and reproducible electronic health record (EHR)-based phenotyping framework for identifying disease cases and generating matched control cohorts for downstream analyses. Here, we developed the Phenotyping Algorithm for Cases and matched Controls using EHR-based Rules (PACER). Materials and Methods Applying PACER to the All of Us Research Program Curated Data Repository v8.0, we identified female breast cancer (BC) cases identified among participants recorded as female at birth using at least two BC-associated diagnostic Observational Medical Outcomes Partnership concept IDs documented at least 30 days apart. A one-to-one matched control cohort was generated by jointly matching on sex, age, genetic ancestry, and state-level residency. Clinical, socioeconomic, and genomic data were integrated for analysis. Results We identified 10,225 BC cases and generated a control cohort of the same size matched for key demographic characteristics. Comparison with a phecodeX-based BC cohort showed 91.03% agreement. Among cases responding to relevant survey items, 80.86% self-reported a personal history of BC, compared to 1.89% of controls. We detected an enrichment of BC-associated GWAS catalog variants, pathogenic mutations in known risk genes, and higher polygenic risk scores in cases compared to controls. Discussion and Conclusion Concordance across a phecodeX-based cohort, self-reported survey responses, and genomic analyses supports the validity of PACER-defined cohorts. PACER is publicly available and readily adaptable to other diseases, supporting future research in risk modeling and precision medicine.

2
Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome

Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.

2026-08-31 oncology 10.64898/2026.08.26.26361432 medRxiv
Top 0.1%
11.9%
Show abstract

Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.

3
Performance of general-population breast cancer risk prediction models in an international consortium

Brantley, K. D.; Ahearn, T. U.; Norton, E. L.; MacInnis, R.; Palmer, J. R.; Fortner, R. T.; Vachon, C. M.; Beane-Freeman, L.; Berrington de Gonzalez, A.; Frost, R.; Bertrand, K. A.; Zirpoli, G.; Neuhouser, M. L.; Barnett, M.; Teras, L. R.; Hodge, J. M.; Patel, A. V.; Bodelon, C.; Lacey, J. V.; Spielfogel, E. S.; Rohan, T. E.; Kirsh, V. A.; Langseth, H.; Tsuruda, K. M.; Milne, R. L.; Haiman, C.; Scott, C. G.; Eliassen, A. H.; Rosner, B.; Willett, W. C.; Romanos-Nanclares, A.; Chen, Y.; Wu, F.; Zheng, W.; Long, J.; O'Brien, K. M.; Sandler, D. P.; Kitahara, C. M.; Linet, M. S.; Anderson, G.; Lars

2026-08-23 epidemiology 10.64898/2026.08.20.26360899 medRxiv
Top 0.1%
11.6%
Show abstract

Background: Several breast cancer (BC) risk prediction models have been developed to provide personal risk assessments. Though individually validated, their performance has not been systematically evaluated across a wide range of populations or ages. Methods: We harmonized individual-level baseline questionnaire data and incident BC diagnoses from 21 cohorts from North America, Europe, and Australia participating in the Breast Cancer Risk Prediction Project. Five-year absolute risk of invasive BC was estimated for five established risk prediction models using classical risk factors only. Discrimination was evaluated by area under the curve (AUC). Calibration was assessed using average and risk-decile specific expected to observed (E/O) ratios. Performance metrics were meta-analyzed across cohorts and models. Metaregression tested associations between cohort characteristics and performance metrics. Results: This analysis included 1,595,977 women aged 20-75 years, enrolled in studies between 1976-2015, with 19,062 (1.2%) invasive BC cases ascertained within 5 years from exposure assessment. Age-adjusted AUCs were similar across models and cohorts (pooled AUCs by model: 0.57-0.58), while E/O ratios varied substantially (pooled E/O ratios by model: 0.83-1.25). Overestimation was common among predicted high-risk individuals (>3%). No appreciable differences in model performance by cohort age, birth year, race, and variable missingness emerged. Calibration improved after assigning race-specific incidence rates. Conclusion: Existing BC risk prediction models provided similar risk discrimination across multiple cohorts, although there was overestimation of risk for high-risk individuals. Performance variation across cohorts was not driven by specific characteristics, which supports development of a unified risk model for diverse populations that leverages appropriate incidence rates.

4
The Role of Distress-related Metabolic Dysfunction in Ovarian Cancer Development: a pooled case-control study

Lin, N.; Balasubramanian, R.; Menichetti, G.; Eliassen, H.; Trabert, B.; Avila-Pacheco, J.; Townsend, M. K.; Terry, K. L.; Clish, C. B.; Tworoger, S. S.; Zeleznik, O. A.

2026-08-31 epidemiology 10.64898/2026.08.27.26361473 medRxiv
Top 0.1%
8.1%
Show abstract

Background: Evidence suggests chronic distress influences ovarian cancer (OC) etiology and metabolomic profiles. Here, we evaluated the association of a metabolite-based distress score (MDS) and OC risk. Methods: We included two matched case-control studies nested within the Nurses' Health Studies (N=584) and the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial (N=348). Metabolites were measured 3-27 years before diagnosis using liquid-chromatography tandem mass spectrometry. We examined the association of quintiles of MDS and 19 constituent metabolites with OC risk using unconditional logistic regression and stratified by tumor histotype, menopausal status, and age at diagnosis. Results: We observed women in the highest versus lowest quintile of MDS had an increased OC risk (OR=1.62,95%CI=1.03-2.54,ptrend=0.07), and type 2 tumors (OR=1.71,95%CI=1.03-2.83,ptrend=0.11). Associations were suggestively stronger for premenopausal and <69-year-old women, and driven by pseudouridine, and N2,N2-dimethylguanosine. Conclusion: Our findings suggest chronic distress-associated metabolic dysregulation may represent a novel OC risk factor, especially among younger women.

5
Cross-Cohort Evaluation of NanoString nCounter Data for Recurrence Prediction in Colorectal Cancer

Quarles Van Ufford, P.; Bojesen, R. D.; Olsen, L. R.; Gogenur, I.; Lund, O.

2026-08-17 oncology 10.64898/2026.08.13.26360359 medRxiv
Top 0.1%
5.5%
Show abstract

Gene expression-based prognostic models have shown promise for predicting recurrence in colorectal cancer (CRC), but their clinical implementation remains limited. The NanoString nCounter platform provides a practical alternative to RNA sequencing and microarrays through standardized, cost-effective gene expression profiling that is compatible with routine clinical samples. In this study, we evaluated whether NanoString nCounter gene expression data improve prediction of recurrence following curative CRC surgery. Gene expression profiles from the NanoString PanCancer IO 360 panel were analyzed in two independent CRC cohorts (cohort A, n = 189; cohort B, n = 131). Differential gene expression analyses and Cox proportional hazards models were used to assess the prognostic value of gene expression alone and in combination with established clinical risk factors. Model performance was evaluated by five-fold cross-validation and external validation between cohorts using the concordance index (C-index) and Kaplan-Meier risk stratification. The two cohorts differed significantly in recurrence-free survival, and differential expression analysis demonstrated marked cohort-specific transcriptional patterns. Ninety-one recurrence-associated genes were identified in cohort A, whereas no significant genes were detected in cohort B, with poor agreement in gene-level differential expression between cohorts (Pearson r = 0.128). Across all prediction models, external performance was modest, and inclusion of gene expression data did not improve prediction beyond clinical variables. The clinical baseline model, incorporating age, UICC stage, and tumor site, consistently achieved the highest cross-cohort performance, with UICC stage emerging as the strongest predictor of recurrence. Although overall discrimination was moderate, the baseline model successfully stratified patients into significantly different high- and low-risk groups across cohorts. These findings indicate that prognostic gene expression signatures derived from NanoString data showed limited reproducibility across independent cohorts and provided little additional predictive value beyond established clinical factors. The results highlight the importance of external validation and suggest that robust clinical variables remain the most reliable predictors of recurrence risk in this setting.

6
Assessing genetic factors, presenting symptoms, and comorbidities in ovarian cancer diagnosis and survival: a retrospective study

Ko, S.; Demirchian, M.; Diaz Miranda, E.; Goldenberg, C.; Krell, K.; Parry, E.; Hunter, M.; Brennaman, L.; Hull, A.; Voth, C.; Lei, L.

2026-08-12 oncology 10.64898/2026.08.11.26360203 medRxiv
Top 0.1%
4.4%
Show abstract

Objective: The purpose of this study is to determine how family history of cancer, genetic mutations, presenting symptoms, and comorbidity burden collectively influence cancer outcomes in patients with epithelial ovarian cancer. Methods: A retrospective analysis was conducted on all patients with epithelial ovarian cancer treated at the University of Missouri and Ellis Fischel Cancer Center between 2008 and 2024. Patient charts were reviewed for histological subtypes, stage of cancer, status of metastasis, CA-125 values, presenting symptoms, comorbidities, family history of cancer, genetic mutations, and survival outcome. Cox regression and association analyses were performed. Results: In this cohort of patients, comorbidities and genetic mutations did not influence ovarian cancer survival. While histological subtypes, CA-125 levels, and cancer stage remained strongly associated with survival. Significant associations were observed between certain presenting symptoms and cancer histological subtype, a family history of breast cancer, stage of cancer at diagnosis, the status of metastasis, and CA-125 levels. Conclusion: Comorbidities and genetic mutations were not significantly associated with ovarian cancer survival. Presenting symptoms were associated with several clinical and pathological variables linked to ovarian cancer diagnosis.

7
Cannabis use and Cancer: Dissecting genetic causality for site-specific risks through two-sample Mendelian Randomization

Lukhere, E.; Kachingwe, B.; Kipandula, W.; Chiphangwi, N.; Singini, M. G.; Kamiza, A. B.

2026-08-13 genetic and genomic medicine 10.64898/2026.08.12.26360176 medRxiv
Top 0.1%
4.3%
Show abstract

Background: The prevalence of cannabis use is increasing at an alarming rate owing to its legalization and decriminalization in some countries. Epidemiological evidence on the association between cannabis use and cancer is inconsistent and conflicting. Herein, we performed two-sample Mendelian randomization (MR) to investigate whether cannabis use is causally associated with site-specific cancers in individuals of European ancestry. Methods: We identified 22 independent genetic variants strongly associated with cannabis use (p-value < 5 x 10-8) in a large meta-analysis of genome-wide association studies of individuals of European ancestry. Genome-wide association summary-level data on site-specific cancers were obtained from individuals of European ancestry in FinnGen, Finland. MR analyses were performed using the inverse-variance weighted (IVW) and multivariable method. Sensitivity analyses were performed using the simple median, weighted median, MR-Egger, and MR pleiotropy residual sum and outlier methods. Results: Our multivariable IVW analyses adjusted for cigarette smoking found that genetic liability to cannabis use was causally associated with esophageal cancer (odds ratio [OR] =1.74, 95% confidence interval [CI] =1.29-2.15, p-value =0.013) and lung cancer (OR=1.35, 95% CI = 1.13-1.58, p-value =0.009). However, genetic liability to cannabis use exerted a protective effect against pancreatic cancer (OR=0.77, 95% CI =0.57-0.91, pvalue=0.032) in individuals of European ancestry in the FinnGen. Our sensitivity analyses found no evidence of horizontal pleiotropy between cannabis use and site-specific cancers. Conclusion: We found that genetic liability to cannabis use was associated with esophageal, lung, and pancreatic cancers in individuals of European ancestry.

8
Site-Specific Cancer Incidence among Clinical Subtypes of Newly Diagnosed Type 2 Diabetes in the United States

Li, Z.; Liu, C.; Weber, M. B.; Ali, M. K.; Hofmeister, C. C.; Varghese, J. S.

2026-08-18 epidemiology 10.64898/2026.08.17.26360595 medRxiv
Top 0.1%
4.1%
Show abstract

Background: Type 2 diabetes (T2D) is associated with elevated rates of several cancers and is increasingly recognized as a heterogeneous disease, but whether its clinically distinct subtypes carry different cancer risks is unknown. Methods: In this matched retrospective cohort study using electronic health record data from the Epic Cosmos Research Platform (2012-2025), adults with newly diagnosed T2D were classified into severe insulin-deficient (SIDD, 21.6%), mild obesity-related (MOD, 23.5%), mild age-related (MARD, 40.7%), or mixed (14.1%) subtypes using validated algorithms and matched to adults without diabetes on age, sex, and body mass index. Cause-specific Cox models estimated adjusted hazard ratios (HRs) for seven site-specific cancers, accounting for competing risks. Cancer screening uptake was assessed as a secondary outcome. Results: Among 575,139 adults with T2D and 689,719 without diabetes (median follow-up, 3.8 years), MARD had the highest cancer incidence (17.3 per 1,000 person-years). Relative to adults without diabetes, rates of colorectal, pancreatic, liver, endometrial, and ovarian cancer were elevated across subtypes, with the highest hazards in SIDD (HR=3.87, 95% CI=3.51 to 4.27) and mixed phenotypes. Prostate cancer rates were lower in all subtypes, most markedly in MOD (HR=0.60, 95% CI=0.55 to 0.64). Rates of breast cancer were higher among mixed (HR=1.12, 95% CI=1.05 to 1.19) and lower among MOD (HR=0.85, 95% CI=0.80 to 0.90). Mammography and prostate-specific antigen screening were lower across subtypes. Conclusions: Site-specific cancer incidence and screening uptake differed across clinically defined subtypes of T2D. Subtype classification from routine clinical data may inform targeted cancer surveillance, though further study is needed before clinical use.

9
Optimal LDCT screening for never-smoking Asian women using integrated polygenic and environmental risk: a microsimulation modelling study

Kowada, A.

2026-08-19 oncology 10.64898/2026.08.18.26360665 medRxiv
Top 0.1%
2.8%
Show abstract

Objective To identify optimal initiation ages and screening intervals for low-dose computed tomography (LDCT) screening among never-smoking Asian women using an integrated polygenic risk score (PRS)-environmental tobacco smoke (ETS) risk model, and to evaluate the cost-effectiveness of alternative screening strategies at these optimized ages. Design Integrated PRS-ETS microsimulation modelling. Setting Japan. Participants Never-smoking women stratified into eight risk groups defined by combinations of PRS levels and ETS exposure. Interventions LDCT screening at intervals of 1 to 10 years, annual chest radiography (CXR), or no screening. Main outcome measures Costs, quality-adjusted life years (QALYs), incremental cost-effectiveness ratios (ICERs), net monetary benefits, lung adenocarcinoma incidence and mortality, and optimal LDCT initiation ages. Sensitivity analyses used a willingness-to-pay threshold of US$50,000 per QALY gained. Results Optimal initiation ages ranged from 40 to 55 years across the eight PRS-ETS risk groups, with higher PRS-ETS risk associated with younger optimal initiation ages. Annual LDCT was the most cost-effective strategy across all PRS-ETS risk strata, yielding an ICER of US$40,471 per QALY in the lowest risk stratum and becoming cost-saving in higher risk strata. Over a lifetime, annual LDCT averted 8,534 lung adenocarcinoma deaths compared with annual CXR and 14,940 deaths compared with no screening. Conclusions Tailoring LDCT initiation age across integrated PRS-ETS risk groups maximizes mortality reduction achievable with cost-effective annual LDCT screening among never-smoking Asian women. These findings highlight an urgent limitation of global lung cancer screening guidelines that rely exclusively on smoking history and provide policy-ready evidence supporting the integration of PRS and ETS into future recommendations for precision LDCT screening for never-smoking populations.

10
Population-scale integration of tumor transcriptomics into breast cancer care: a decade of the SCAN-B initiative

Saal, L. H.; Dalal, H.; Meng, P.; Brueffer, C.; Gladchuk, S.; Gruvberger-Saal, S. K.; Hakkinen, J.; Nordborg, N.; Li, M.; Valcich, J.; Hedenfalk, I.; Edsjo, A.; Killander, F.; Nimeus, E.; Bendahl, P.-O.; Forsare, C.; Manjer, J.; Malina, J.; Rehn, M.; Ahsberg, K.; Ingvar, C.; Graffner, F.; Ahlund, L.; Asking, B.; Erngrund, M.; Sjovall, M.; Cetti, A.; Svensjo, T.; Teder, H.; Bjorkman, J.; Myrskog, L.; Falck, A.-K.; Kallstrom, A.-C.; Einebigi, Z.; Braganca, P. R.; Lindman, H.; Sjoblom, T.; Malmberg, M.; Larsson, C.; Ehinger, A.; Ryden, L.; Loman, N.; Hegardt, C.; Borg, A.; Vallon-Christersson, J.

2026-08-23 oncology 10.64898/2026.08.20.26360879 medRxiv
Top 0.1%
2.6%
Show abstract

Background: Population-scale molecular profiling integrated into routine healthcare could accelerate biomarker discovery, validation, and implementation, but the feasibility and sustainability of such an approach have rarely been demonstrated prospectively. The Sweden Cancerome Analysis Network - Breast (SCAN-B) Initiative was established to integrate prospective molecular profiling with population-based breast cancer care and create an infrastructure for translating molecular discoveries into clinical practice (ClinicalTrials.gov identifier NCT02306096). Methods: We evaluated the first 10 full calendar years of SCAN-B, encompassing patients with primary invasive breast cancer enrolled between August 30, 2010 and December 31, 2020. Enrollment and biospecimen collection were compared with all eligible breast cancer diagnoses in participating hospitals to assess population coverage and representativeness. Clinicopathological characteristics, treatments, recurrence-free survival, overall survival, RNA-sequencing-based molecular subtypes and risk-of-recurrence, and somatic mutations were evaluated. We additionally report the translation of SCAN-B molecular profiling from the research setting into routine clinical diagnostics. Results: Among 16,381 estimated eligible breast cancer diagnoses, 13,940 patients (85.1%) were prospectively enrolled across participating Swedish hospitals. Baseline blood samples were obtained from 98.4% of enrolled patients and tumor specimens from 71.1%; 9,323 tumors (94.0% of submitted tumor specimens) underwent RNA-sequencing. The enrolled cohort was broadly representative of the underlying breast cancer population across major clinicopathological characteristics. Integration of longitudinal clinical data with molecular profiling enabled characterization of real-world treatment patterns, long-term outcomes, molecular subtypes, risk-of-recurrence, and the somatic mutational landscape in this population-based cohort. Building on prospective real-time RNA-sequencing and subsequent development and validation of single-sample molecular subtype and risk-of-recurrence predictors, the SCAN-B workflow was transferred into routine clinical molecular diagnostics in Sk[a]ne and Blekinge in 2021. Through January 2026, more than 3,000 patients had received clinical RNA-sequencing-based molecular subtype and risk-of-recurrence reports, while prospective SCAN-B enrollment and transfer of samples and molecular data into the research infrastructure continued. Patient enrollment continues prospectively, with over 23,000 patients accrued as of January 2026. Conclusions: A prospective, population-based molecular profiling program can be integrated into routine breast cancer care at scale while maintaining high population coverage and representativeness. Over more than a decade, SCAN-B progressed from prospective biosampling and molecular profiling through biomarker development and validation to implementation of RNA sequencing-based testing in routine healthcare. This model establishes a continuous framework linking population-based molecular research, biomarker discovery and validation, and clinical implementation, and provides a strategy for integrating precision oncology research with routine cancer care.

11
Assessing ethnic differences in age-standardised net survival of eight common cancer: an English population-based study

Martins, T. O.; Rachet, B.; Hamilton, W.; Majano, S. B.

2026-08-07 epidemiology 10.64898/2026.08.05.26359765 medRxiv
Top 0.1%
2.5%
Show abstract

Background: We examined ethnic differences in age-standardised net survival (ANS) for eight common cancers diagnosed in England between 2010 and 2019. Methods: Analyses included 247,428 patients aged [&ge;]40 years diagnosed with breast, prostate, lung, colorectal, cervical, ovarian, myeloma, and oesophagogastric cancers. Net survival was estimated at one, three, and five years using the Pohar-Perme estimator and age-standardised with International Cancer Survival Standards weights across four age bands. Results: Compared with White patients, Black patients had higher ANS for lung and prostate cancers at all time points, for myeloma at one year, and for oesophagogastric cancer at one and three years. However, they had lower ANS for breast cancer at three years. Asian patients had higher ANS for lung, prostate, and oesophagogastric cancers at all time points, and for other sites at varying follow-up times. Patients in the Mixed group had higher ANS for most cancers, whereas those in the Other ethnic group generally had lower ANS compared with White patients. Conclusions: Ethnic minority groups in England do not consistently experience poorer cancer survival, with varying patterns observed by cancer site. Universal healthcare access may reduce disparities observed elsewhere, highlighting the importance of context-specific research and public policy.

12
New-onset type 2 diabetes mellitus and obesity-related cancer risk: a matched cohort study (UK Biobank)

Tipping, O.; Wang, M.; Martin, R.; Sperrin, M.; Renehan, A.

2026-08-27 epidemiology 10.64898/2026.08.25.26360729 medRxiv
Top 0.1%
2.4%
Show abstract

Background: Observational research reports positive associations between type 2 diabetes mellitus (T2DM) and obesity-related cancers (ORCs), but causality remains unclear due to confounding (namely the shared risk factor of obesity, commonly approximated as body mass index, BMI), immortal time bias, and detection-time bias. Here, we aimed to use causal inference methods to minimise the above problems and estimate causal associations between new-onset T2DM and incident cancer. Methods: We performed a cohort study within UK Biobank, comparing new-onset T2DM with unexposed individuals matched 1 to 3 on BMI, age, and sex using a sequential longitudinal approach. The primary outcomes were total incident cancer, divided into ORCs and non-obesity-related cancers (NORCs). The secondary outcomes were site-specific cancers. We developed Cox models to estimate time-split hazard ratios (tsHRs) and 95% confidence intervals (CIs) stratified by sex. Findings: 23,771 participants with new-onset T2DM were matched with 71,170 unexposed participants. During a median follow-up of 5 years, there were 7694 (T2DM: 2432; unexposed: 5262) incident cancers. In men, there was evidence for an effect of T2DM on obesity-related cancer (tsHR 1.39, 95% CI 1.21-1.59), particularly on hepatocellular carcinoma (tsHR 3.97, 95% CI 2.38-6.65), pancreatic (tsHR 1.77, 95% CI 1.15-2.72) and kidney (tsHR 1.62, 95% CI 1.13-2.32) cancers. In women, there was evidence for an effect on obesity-related cancers (tsHR 1.33, 95% CI 1.16-1.52). Importantly, there were no associations with post-menopausal breast and endometrial cancers, two cancer types consistently associated with elevated BMI. There was no effect of new-onset T2DM on incidence of NORCs. There was evidence of detection-time bias, particularly in men. Interpretation: This is the first large-scale study to demonstrate evidence of a BMI-independent associations between new-onset T2DM and incident cancer. In men, this was primarily driven by hepatocellular carcinoma, pancreatic cancer, and kidney cancer. In women, the underlying cancers driving this relationship were less clearly defined. Funding: This study was funded by Cancer Research UK and administered through the Manchester Cancer Research Centre MB-PhD scheme (SEBCATP-2023/100010).

13
A 12-year retrospective analysis of cancer epidemiology at a major oncology center in Erbil, Iraq

Marouf, S. S.

2026-08-10 oncology 10.64898/2026.08.05.26359839 medRxiv
Top 0.1%
2.2%
Show abstract

Long-term data describing cancer patterns in Iraq remain generally limited. This retrospective observational study was conducted to evaluate the distribution and longitudinal patterns of malignant solid tumors diagnosed over a period of 12 years (2014-2025) at a major tertiary oncology center in the Kurdistan Region of Iraq. Demographic characteristics, cancer types, and temporal trend changes in cancer distribution were analyzed. Comparisons were made between the first (2014-2019) and second (2020-2025) halves of the study period. Descriptive statistics, Chi-square tests, and linear regression analyses were used to evaluate the temporal trends. After excluding records with incomplete data, 11,704 patients were included. Breast cancer was the most frequently diagnosed malignancy, accounting for over one-quarter of all cases, followed by lung, colorectal, prostate, and bladder cancers, in descending order. Women represented the majority of patients, and the mean age at diagnosis increased significantly over time. In general, the relative distribution of colorectal, genitourinary, pancreatic, uterine, and thyroid cancers increased during the study period, whereas breast and lung cancers showed a modest but significant proportional decline despite remaining the most common malignancies. A marked reduction in case numbers was observed in 2020, followed by progressive recovery in subsequent years. To conclude, cancer patterns in the Kurdistan Region of Iraq changed substantially during the 12-year study period, with increasing proportions of colorectal and several other malignancies alongside an older age at diagnosis. These findings likely reflect a combination of demographic changes, evolving lifestyle-related risk factors, and improvements in cancer detection and referral. The results provide contemporary evidence to support regional cancer control strategies, screening programs, resource allocation, and future epidemiological research.

14
Loss of NKX2-1 predisposes thyroid to neoplasm development through regulation of oxidative stress

Shirai, Y.-T.; Ward, J. M.; Takizawa, Y.; Liu, H.; Miyakoshi, M.; Iwadate, M.; Murata, T.; Hayase, S.; Yokoyama, S.; Ehata, S.; Kimura, S.

2026-08-27 cancer biology 10.64898/2026.08.26.746581 medRxiv
Top 0.1%
2.2%
Show abstract

Many factors including ionizing radiation and iodine deficiency are known to increase thyroid carcinogenesis risk. Our dataset analysis of The Cancer Genome Atlas (TCGA) showed that lower mRNA expression of NK2 homeobox 1 (NKX2-1) transcription factor, a master regulator of genesis, homeostasis, and function of thyroid, is linked to poor prognosis of papillary thyroid cancer patients. Here we provide the findings that thyroid-specific Nkx2-1 conditional knockout (Nkx2-1{Delta}T) mice develop thyroid adenoma and carcinoma in higher frequency with combined exposure to radiation and iodine deficiency than control Nkx2-1fl/fl mice. Iodine deficiency caused oxidative stress, which subsequently resulted in DNA damage, leading to transformation of thyroid follicular cells. RNA-seq gene set enrichment analysis indicated higher production of reactive oxygen species (ROS) in the thyroids of Nkx2-1{Delta}T as compared to Nkx2-1fl/fl mice with combined exposure to radiation and iodine deficiency. This was accompanied by a feedback induction of SOD3 (superoxide dismutase 3) and GPX2 (glutathione peroxidase 2). These antioxidants were naturally expressed at higher levels in the thyroids of Nkx2-1{Delta}T than Nkx2-1fl/fl mice without iodine deficiency or radiation. Nkx2-1{Delta}T thyroids exhibited abnormal follicle architecture and up-regulation of Acox2 (encoding acyl-CoA oxidase 2), which produces hydrogen peroxide. These results suggest that loss of NKX2-1 may contribute to excess ROS production, which elevates basal oxidative stress resulting in the promotion of ROS-induced carcinogenesis. We propose a role for NKX2-1 as a regulator of ROS production homeostasis in the thyroid. Its disturbance would dispose thyroid follicular cells more vulnerable to the ROS-producing carcinogens.

15
The impact of neighborhood socioeconomic deprivation on metastatic pancreatic cancer treatment and survival: An incidence-based, causally-structured observational study

Raghu, A.; Shah, S.; Pattnaik, A.; Permuth, J. B.; Park, M. A.; Dhahri, H.; Huang, H. C.; Fleming, J. B.; Anaya, D. A.; Powers, B. D.

2026-08-10 oncology 10.64898/2026.08.06.26359821 medRxiv
Top 0.1%
2.2%
Show abstract

Purpose: Metastatic pancreatic ductal adenocarcinoma (PDAC) portends a poor prognosis. Prior studies have assessed the association of socioeconomic deprivation (SED) in PDAC often with large geographic areas. This study employed a causal framework to characterize neighborhood SED on treatment receipt and survival in metastatic PDAC. Methods: Using the incidence-based Florida Cancer Data System, metastatic PDAC patients diagnosed from 2007-2015 were identified. The Area Deprivation Index, a composite measure of SED that ranks neighborhoods from 1-100 (higher scores = higher deprivation), was used to assess receipt of systemic therapy and overall survival (OS). Exposures and covariates were assessed using descriptive statistics and a causal inference framework. Results: Overall, 9,574 patients met inclusion criteria. 46.6% of patients received systemic therapy, ranging 39.4% to 54% in the highest and lowest SED quartiles, respectively. After adjustment, the lowest quartile had increased odds of systemic therapy relative to the highest (OR 1.93; 95% CI 1.70-2.18). Median OS was 3.8 months for the lowest quartile and 2.4 months for the highest (p = 0.01). Patients in the highest quartile had an estimated 32% higher hazard of death than the lowest (HR 1.32, 95% bootstrap CI 1.20-1.40). Conclusion: In an incidence-based statewide cohort, most patients did not receive treatment for metastatic PDAC and median OS was poor-2.9 months. Using a causal inference framework, higher SED led to lower rates of systemic therapy receipt and worse overall survival in metastatic PDAC. Future research should focus on the mechanisms that shape these findings.

16
A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy

Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26360896 medRxiv
Top 0.2%
1.7%
Show abstract

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

17
Trends in the Utilization of Breast, Cervical, and Colorectal Cancer Screening from 2010 to 2019 Among a Commercially Insured Population Using the MarketScan Commercial Claims Database

Sun, J.; Wat, R.; Frick, K. D.; Kong, X.; Liang, H.; Chow, C.; Shi, L.

2026-08-11 epidemiology 10.64898/2026.08.09.26360037 medRxiv
Top 0.2%
1.5%
Show abstract

Introduction: Breast, cervical, and colorectal cancer screening guidelines changed substantially between 2010 and 2019. We examined trends in the annual utilization of these screenings among commercially insured enrollees in the United States from 2010 to 2019 by age group, geographic region, and screening modality. Methods: We conducted a retrospective, serial cross-sectional analysis of the MarketScan Commercial Claims Database from 2010 through 2019, comprising approximately 141.2 million privately insured enrollees. Annual screening rates, defined as the proportion of eligible enrollees receiving a given test within each calendar year, were estimated for cervical, breast, and colorectal cancer using procedure codes, stratified by age group, screening modality, and geographic residence. These reflect annual utilization rather than up-to-date (guideline-concordant) screening. Temporal trends were evaluated using two-sided Poisson regression, and urban-rural disparities in 2019 were assessed using multivariate generalized estimating equations. Results: Cancer screening utilization remained stagnant or declined across all three cancer types over the study period. Among women aged 30-64 years, cervical cytology alone declined substantially from 28.2% in 2010 to 8.8% in 2019, while co-testing increased from 11.4% to 20.3%. Screening mammography among women aged 50-64 showed minimal change, remaining stable at 45.7% in 2010 and 45.8% in 2019. Colorectal cancer screening across enrollees aged <64 decreased modestly from 7.7% in 2010 to 6.5% in 2019, with a more pronounced decline among adults aged 45-49 years. Across all three cancer types, screening utilization was higher among urban residents than rural residents, with incidence rate ratios ranging from 1.02 to 1.05 in 2019. Conclusions: Utilization of cervical, breast, and colorectal cancer screening among commercially insured adults did not improve between 2010 and 2019. Persistent urban-rural disparities highlight ongoing gaps in preventive care delivery. Targeted interventions may help improve screening utilization, particularly in rural and underserved populations.

18
Readability Assessment of Patient-Reported Measures Used During Heritable Cancer Genetic Testing

Adegbesan, A. C.; FitzGerald, L.; Dickinson, J. L.; Raspin, K.; Roydhouse, J.

2026-08-17 oncology 10.64898/2026.08.13.26360322 medRxiv
Top 0.2%
1.2%
Show abstract

Background: Patient-reported measures (PRMs), including patient-reported outcome and experience measures, capture patients perspectives on their health status and healthcare experiences. In cancer genetics, PRMs have been used to assess genetic knowledge, psychosocial outcomes, and decision-making. However, patients must understand these measures to provide useful information, an ability which is influenced by general and health literacy levels. Readability guidelines recommend that patient-facing materials be written at or below a Grade 6 level. This study evaluated the readability of PRMs used in a cancer genetic testing context. Objective: To assess whether PRMs used in heritable cancer genetic testing meet recommended readability levels using validated indices. Methods: PRMs were identified from a recent systematic review of PRMs used in heritable cancer genetic testing, which reported 83 instruments across eight categories. English-language PRMs containing structured question items and response scales were eligible for extraction and converted into plain text for analysis. Readability was assessed using four validated indices: Flesch Kincaid Grading Level (FKGL), FORd, CAylor, and STicht (FORCAST) formula, Flesch Reading Ease Score (FRES), and Simple Measure of Gobbledygook (SMOG) via an automated readability software. Descriptive analysis and numerical comparison evaluated readability levels across PRM categories and against the recommended Grade 6 reading level. Results: Sixty-five PRMs met the eligibility criteria, with most, including validated instruments, exceeding the recommended Grade 6 reading level. Across the eight categories, genetics-specific PRMs required the highest readability levels, indicating higher readability demands. Conclusions: Most PRMs, particularly those specific to genetics, do not meet readability guidelines. This may limit their accessibility to individuals with limited general and health literacy. Development of PRMs specific to genetics should consider strategies to improve readability, such as plain-language approaches and involvement of individuals with limited general or health literacy. Keywords: readability, patient-reported measures, cancer, genetic testing, health literacy

19
Genetic prediction of colorectal cancer risk in six major ancestries provides insights to streamline practice screening guidelines.

Parasuraman, A.; Lim, A. W.-Y.; Eltayib, R.; Pandeya, N.; Olsen, C. M.; Radford-Smith, G.; Whiteman, D. C.; MacGregor, S.; Seviiri, M.

2026-08-11 gastroenterology 10.64898/2026.08.10.26360074 medRxiv
Top 0.2%
1.1%
Show abstract

Background and objective: Colorectal cancer (CRC) is the third leading cause of cancer deaths worldwide. Early identification of high-risk individuals allows targeted prevention and early detection. Design: We constructed a polygenic risk score (PRS) for CRC risk using data from 1,448,354 individuals (103,401 cases). We evaluated its performance for identifying high-risk individuals in 6 major ancestries. Results: The PRS was strongly associated with CRC risk in Europeans (OR per SD =2.13, 95%CI=1.98-2.28), Africans (OR=1.35, 95%CI=1.11-1.64), Hispanics (OR=1.97, 95%CI =1.48-2.61), East Asians (OR=1.98, 95%CI=1.35-2.91), South Asians (OR=1.85, 95%CI=1.44 -2.37), and Middle Easterners (OR=3.10, 95%CI=1.38-6.95). Europeans in the top 10% genetic risk had 14-fold and 5-fold higher CRC risks compared to the bottom 10% (OR=13.50, 95%CI=8.67-21.00), and average (20-70%) risk groups (OR=4.64, 95%CI=3.89-5.52), respectively. The CRC risk in the top 10% individuals was equivalent to having three affected first degree relatives with CRC diagnosed at any age. Genetically high-risk individuals developed CRC up to 15 years earlier than the average. The PRS was strongly associated with early onset CRC risk e.g. in AFR (OR=3.22, 95%CI=1.79-5.81), and improved its prediction e.g. by 9% beyond clinical predictors in EUR. Conclusion: A comprehensive genetic prediction of CRC risk provides insights that could streamline screening and prevention guidelines.

20
Potential Cost of Multi-Cancer Detection Tests to Medicare

Scherer, L. D.; Matlock, D. D.; Cronin, J.; Gritz, M.

2026-08-27 health economics 10.64898/2026.08.24.26361257 medRxiv
Top 0.2%
1.1%
Show abstract

Multi-Cancer Detection (MCD) tests can detect more than 50 different types of cancer using a blood test. Recently passed law in the U.S. guarantees that Medicare will pay for these tests when they are FDA approved and show evidence for clinical benefit. This manuscript provides estimates of the cost of MCD tests to Medicare under different assumptions of cost per test, eligibility, and screening uptake in the eligible population. This manuscript additionally estimates the cost of follow-up testing resulting from false positive results, which are considered avoidable costs caused by the screening test.